Reinforcement Learning Startups funded by Y Combinator (YC) 2026

September 2026

Browse 40 of the top Reinforcement Learning startups funded by Y Combinator.

We also have a Startup Directory where you can search through over 5,000 companies.

  • Magma
    Magma
    Y Combinator LogoS2026
    Active • 1 employees • San Francisco
    We enable companies monetize their agents' traces.
    ai
    reinforcement-learning
    data-engineering
  • EdotEnv
    EdotEnv
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    A Quant Neolab building toward RSI by leveraging markets as self-improving research environments. They work with frontier AI labs and academic groups building research harnesses, evaluation benchmarks, and post-training environments.
    time-series
    reinforcement-learning
  • Enact
    Enact
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Enact creates the physical environments to benchmark and hill climb robotics models on high-value tasks. We find where robots fail, and fix them with our in-house training loop.
    reinforcement-learning
    robotics
    ai
  • Prodigy Research
    Prodigy Research
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Prodigy is the frontier AI trading research lab. We've trained the world's best foundation model for quantitative finance. We've already achieved more than 100% returns in live trading over the course of our YC batch, doubling our capital while major indices were flat. Our founders previously worked on a trading desk at Jane Street, trained frontier AI models at Google DeepMind and shipped cutting-edge AI at Apple to billions of customers. Join us as we use AI to solve the market.
    investing
    reinforcement-learning
    ai
    trading
    finance
  • Riften
    Riften
    Y Combinator LogoS2026
    Active • 5 employees • San Francisco
    Intelligence will be the most valuable asset a company owns. Today, almost every company rents it. Companies pay a handful of labs one token at a time. Each dollar buys an answer, not an asset. They spend more, but own nothing more. Riften reverses that. Change one environment variable and Riften becomes the gateway for a company's AI traffic. We route each request to the lowest-cost model that can do the job, serving most work on open-weight models we host. Customers save money immediately without rebuilding their products. The gateway is the entry point. Over time, Riften turns the company's work into private models built on open weights and shaped around the business. What begins as cheaper inference becomes an intelligence stack the company owns, runs on infrastructure it controls, and uses to power its most important applications. We call this earned intelligence. It is created by the company's own work, controlled by the company, and accumulated over time. The next generation of companies will not rent their intelligence. They will earn it, own it, and compound it. Riften is how they get there.
    developer-tools
    open-source
    reinforcement-learning
    infrastructure
    ai
  • hiloop
    hiloop
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    hiloop helps teams train agents for tasks where general models are not good enough. Give us a task, your current agent or model, and an evaluation. hiloop runs an autoresearch campaign across data, SFT and other post-training methods, continual learning, prompts, tools, harnesses, and systems, then returns the best verified improvement. It runs hosted or in your cloud. We provide the research system around models: persistent memory, full experiment lineage, compute orchestration, and statistical verification. We’re starting with agent and model training, continual learning, and optimization. Reach out to us for early access at founders@hiloop.ai.
    developer-tools
    infrastructure
    machine-learning
    reinforcement-learning
    artificial-intelligence
  • Olam Labs
    Olam Labs
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Multi-agent environments let us evaluate and train models in complex, simulated worlds. We work with researchers on both evaluations and training for character and agentic performance typically difficult to assess for using current popular datasets. Play one of our first releases, Multi-Agent Arena (https://olamlabs.ai/arena), where humans come play social strategy games against multiple AI agents. It's currently top 50 on OpenRouter and has thousands of matches played each day. We use Multi-Agent Arena to build datasets on agentic performance and real-world socialization.
    reinforcement-learning
    gaming
    data-engineering
    artificial-intelligence
  • Ooak Data
    Ooak Data
    Y Combinator LogoS2026
    Active • 5 employees • San Francisco
    Ooak Data turns internal company data into a training ground where AI agents can be tested on real work.
    reinforcement-learning
    artificial-intelligence
  • Fabraix
    Fabraix
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Fabraix builds state-of-the-art AI red-teaming agents that continuously detect security vulnerabilities in customer-facing AI. Our product, Nyx, has already found vulnerabilities in agents at dozens of Fortune 500 companies. On AgentHarm, the leading benchmark for offensive AI security, Nyx achieved a 78% attack success rate, compared with 67% for GPT-5.6 Sol. AI agents can change more often than teams can test them. A new model, prompt, tool, permission, or data source can change what an agent does, even when the application code stays the same. And AI is also increasing how much software companies produce and how often it changes. We built Nyx to automate the work required to red-team AI agents. It is multi modal by design and interacts with the target indirectly through controlled replicas of SaaS products and websites that we maintain, placing malicious payloads in webpages, documents, files, messages, and tool outputs to test how an agent behaves under different conditions. Nyx often finds its first vulnerability within minutes or hours rather than days or weeks. Because the work is fully automated, companies can repeat the test with every change and at a much lower cost than what a 7-figure red-teaming (and point-in-time) engagement would require. We're also the team behind ACE (Adversarial Cost to Exploit), a benchmark that measures AI security in terms of how much it costs attackers to break an AI system; providing a game-theoretic framework to understand how motivated a rational attacker would be in exploiting the system.
    cybersecurity
    reinforcement-learning
    artificial-intelligence
  • Praxis Robotics
    Praxis Robotics
    Y Combinator LogoS2026
    Active • 3 employees • San Francisco
    Praxis captures the data physical AI is bottlenecked on: egocentric video, 3D scans, and multimodal capture from inside real industrial and residential environments. Embedded within publicly listed and unicorn-scale conglomerates, we reach 60k workers across 5 continents and 150+ environment types, supplying the real-world training data frontier labs and humanoid companies can't source anywhere else.
    reinforcement-learning
    hard-tech
    robotic-process-automation
  • Experiential Labs
    Experiential Labs
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    We are building an open source gateway for AI budget holders to control model access and turn their traffic into better models that they own. We're a team that has earned an ML PhD at Georgia Tech, scaled AutoGPT to 160k GitHub stars, and led a self-driving simulation team at Waabi.
    reinforcement-learning
    infrastructure
    artificial-intelligence
  • CoArena
    CoArena
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    CoArena is a live arena where anyone can use the world's top computer-use models racing two of them on the same real computer task and judging which one did it better. Every battle becomes something the AI labs can't build themselves, an honest test of their agents on real work, and the data to make them better.
    aiops
    reinforcement-learning
    data-labeling
    artificial-intelligence
    machine-learning
  • Markov
    Markov
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Markov builds expert computer-use datasets for frontier AI labs. Our datasets include screen recordings + mouse/keyboard actions of professionals using CAD, design and enterprise software. We work with the top AI labs and we’ve gotten more than 350k downloads on HuggingFace. Our long-term goal is to be the full stack platform for computer-use AI, starting with data and then the infrastructure to train and deploy computer-use models. Dev and Harish are technical co-founders. Dev studied aerospace engineering at IIT Madras and was the youngest team member at Sarvam - India’s top AI lab. Harish is an Emergent Ventures grantee and his prior robotics work was featured in top national publications.
    reinforcement-learning
    data-labeling
    data-engineering
  • Anchorhead
    Anchorhead
    Y Combinator LogoW2026
    Active • 1 employees • San Francisco
    We build custom evals for product teams, as a service. To build these eval sets, we source domain experts and have them turn their work into original, real world evals. For example, our research engineering experts have create tasks involving optimizing an algorithm, deploying a model, or running experiments to solve a novel problem. A grader scores the agent’s performance, and these scores serve as signals during evaluations.
    reinforcement-learning
  • Ashr
    Ashr
    Y Combinator LogoW2026
    Active • 2 employees • San Francisco
    Ashr Manifold is a self-contained model training, hosting, observability, and continual learning platform for purpose-built open-weight models. Startups can use Manifold to post-train and manage fleets of continuously improving open-weight models without having a dedicated research team or infrastructure. Ultimately, this reduces operating costs and allows a greater degree of control over proprietary company knowledge. Enterprises can use Manifold to encode institutional judgements and priorities into open-weight models, reduce token costs, and manage every level of the AI life cycle, from data curation, to model serving on weights and infrastructure they own.
    artificial-intelligence
    reinforcement-learning
    data-engineering
    generative-ai
    deep-learning
  • Brumby (Formerly GrazeMate)
    Brumby (Formerly GrazeMate)
    Y Combinator LogoW2026
    Active • 3 employees • Sydney NSW, Australia
    Brumby builds autonomous drones that herd cattle. On command, our drones fly to a paddock, position themselves around the mob, and move them where they need to go. What used to take a full day of helicopters, motorbikes, and horses now runs on a schedule. We work with some of the largest cattle ranches in the world. While the drones are herding, they're also estimating animal weights, measuring grass biomass, monitoring water levels, and flagging sick animals. We're building physical AI that lets a grazier manage thousands of head across millions of acres from their phone.
    agriculture
    reinforcement-learning
    computer-vision
    drones
  • Cortex AI
    Cortex AI
    Y Combinator LogoF2025
    Active • 3 employees • San Francisco
    Cortex AI builds the world’s most diverse and large-scale real-world workplace robot & egocentric dataset — where the physical world becomes the next training and evaluation set for embodied AI. We power frontier labs developing robotics foundation models and general-purpose robots by providing the data they need: 1️⃣ Egocentric Data — real-workplace human video with hand/body pose, depth, and subtask labels. 2️⃣ Robot Data — trajectories collected from manipulators and humanoids in real industry settings. 3️⃣ Human-in-the-Loop Rollouts & Evals — real-world deployments with remote operators who recover robots when they fail, capturing data that feeds back into training and continuously improves models. Additionally, through the Cortex Marketplace, workplaces get paid to host data-collection and evaluation sessions, while labs access the in-the-wild data that truly matters. This draws on Lucas’s previous experience as co-founder of Carousell, a C2C marketplace that scaled to a $1B+ valuation.
    robotics
    reinforcement-learning
    artificial-intelligence
  • hillclimb
    hillclimb
    Y Combinator LogoF2025
    Active • 4 employees • San Francisco
    We work with frontier AI labs to help train their agents to become AI research scientists
    reinforcement-learning
  • Topological
    Topological
    Y Combinator LogoS2025
    Active • 2 employees • San Francisco
    Topological is developing physics-based foundation models for CAD optimization. We help hardware teams iterate at the same speed that software teams do. Our technology is accelerating the engineering workflow with AI and scales design and optimization to identify the ideal designs for complex problems given their physical constraints with enhanced speed and performance. Our first model, UToP-v1, is a SOTA topology optimization model that understands physics, geometry, and manufacturability. It can generate the most efficient design given a problem’s physical requirements. It has <5% compliance error and is 1930x faster than current methods. We're reimagining mechanical engineering and computational design with precision spatial AI.
    ai
    3d-printing
    design-tools
    reinforcement-learning
    robotics
  • Monte
    Monte
    Y Combinator LogoS2025
    Active • 3 employees • San Francisco
    Monte is an applied research lab building products for continual learning and recursive self-improvement for AI agents.
    b2b
    infrastructure
    reinforcement-learning
    ai
  • idler
    idler
    Y Combinator LogoS2025
    Active • 12 employees • San Francisco
    idler is a frontier data research lab. We build the evals and environments that the world's leading frontier labs use to measure and train their models.
    reinforcement-learning
  • Kairos
    Kairos
    Y Combinator LogoP2025
    Active • 3 employees • San Francisco
    Kairos closes the last-mile reliability gap in AI deployments. We bring frontier techniques to companies in critical industries, deploying specialized agents that encode operator expertise and reliably automate their most manual workflows.
    reinforcement-learning
    aiops
    artificial-intelligence
    machine-learning
  • Aviro
    Aviro
    Y Combinator LogoP2025
    Active • 2 employees • New York City
    We design challenging worlds where AI models learn to explore and solve problems. We turn these worlds into post-training datasets and reinforcement learning environments to train the next frontier of AI. We're research partners with four top frontier AI labs, several Fortune 100 companies, and leading RL data vendors.
    reinforcement-learning
  • Cartpole
    Cartpole
    Y Combinator LogoP2025
    Active • 2 employees • San Francisco
    We're creating reinforcement learning environments for training frontier models.
    reinforcement-learning
    ml
    artificial-intelligence
    data-labeling
  • Freesolo
    Freesolo
    Y Combinator LogoP2025
    Active • 4 employees • San Francisco
    Freesolo works with companies to encode user-trajectory knowledge into specialized models that outperform the state of the art at much lower latency and cost.
    reinforcement-learning
    b2b
    ai
  • HUD
    HUD
    Y Combinator LogoW2025
    Active • 15 employees • San Francisco
    HUD is the platform for building high quality post training datasets. Over 50 businesses use HUD to build RL environments, sell them to AI labs, or train their own models from them. Our mission is to enable a generation of data entrepreneurs. The previous generation built apps to impact the world. We believe people will build infrastructure around data, both digital and physical, that align AIs to their specific goals.
    artificial-intelligence
    reinforcement-learning
  • Agentin AI
    Agentin AI
    Y Combinator LogoW2025
    Active • 2 employees • San Francisco
    At Agentin AI, we build Agents that move data and take actions across enterprise systems, like Salesforce, NetSuite and SAP. These agents are difficult to build because each enterprise heavily customizes their systems but we solved that by training our Agents to learn and adapt from failures, applying reinforcement learning techniques we developed.
    enterprise
    reinforcement-learning
    ai
  • TrainLoop
    TrainLoop
    Y Combinator LogoW2025
    Active • 6 employees • San Francisco
    TrainLoop makes it effortless for developers to supercharge LLM performance through reinforcement learning.
    developer-tools
    generative-ai
    reinforcement-learning
  • Osmosis
    Osmosis
    Y Combinator LogoW2025
    Active • 6 employees • San Francisco
    Osmosis is a post-training platform that helps companies fine-tune language models using reinforcement learning. We work with fast-growing AI companies to train task/domain-specific models that beat foundation models on performance, cost, and latency. Our platform handles compute orchestration, reward modeling, and training run observability as a CLI-based product usable by developers and agents.
    reinforcement-learning
    machine-learning
    infrastructure
    artificial-intelligence
  • Synth
    Synth
    Y Combinator LogoF2024
    Active • 2 employees • San Francisco
    Choose a coding agent harness, model, and task dataset and optimize context and prompts to get the best performance for long-horizon tasks
    ai
    reinforcement-learning
  • Vibrant Labs
    Vibrant Labs
    Y Combinator LogoW2024
    Active • 6 employees • San Francisco
    We work on methods to autonomously scaling evals/environments for post-training agents.
    generative-ai
    open-source
    developer-tools
    ai
    reinforcement-learning
  • JustAI
    JustAI
    Y Combinator LogoW2024
    Active • 4 employees • San Francisco
    Always-on AI agents for 1-1 personalization at scale
    marketing
    personalization
    reinforcement-learning
    workflow-automation
    artificial-intelligence
  • Velos
    Velos
    Y Combinator LogoW2023
    Active • 3 employees • San Francisco
    Velos helps non-technical operations teams automate complex, manual back-office tasks with AI workers instead of overseas teams. Unlike traditional robotic process automation (RPA) platforms like UiPath, Velos automations use machine learning to reliably handle ambiguities in their tasks, eliminating the need for an army of maintenance engineers and consultants to build and maintain your automations. We're automating the repetitive work people hate to do.
    generative-ai
    reinforcement-learning
    artificial-intelligence
    automation
    robotic-process-automation
  • Unify
    Unify
    Y Combinator LogoW2023
    Active • 4 employees • London
    Building the next frontier of models that continually learn from a stream of experience.
    ai
    artificial-intelligence
    open-source
    reinforcement-learning
    infrastructure
  • rct AI
    rct AI
    Y Combinator LogoW2019
    Active • 40 employees • Los Angeles
    rct AI is providing AI solutions to the game industry and building the true Metaverse with AI generated content. By using cutting-edge technologies, especially deep learning and reinforcement learning, rct AI creates a truly dynamic and intelligent user experience both on the consumers’ side and production’s side. The founding team ever built a company, Raventech together and helped make it acquired by Baidu (NASDAQ:BIDU) in 2017.
    gaming
    reinforcement-learning
    metaverse
  • Sepal AI
    Sepal AI
    Y Combinator LogoS2024
    Acquired • 15 employees • San Francisco
    Sepal is a data research company on a mission to advance human knowledge and capabilities through safe AI. We partner with the world’s leading AI labs and enterprises to help their models get better at the tasks people actually want them to do. We’ve built a Cloud-Native Agent Dataset Factory which turns the process of generating evaluation and training data from manual, inconsistent, and labor-intensive into something automated, standardized, and scalable. Sepal AI was founded in 2024 by engineers and operators from Vercel and Turing. We went through Y Combinator, raised several million dollars from leading investors, and already count multiple Fortune 500s and top AI research labs as paying customers.
    data-labeling
    aiops
    reinforcement-learning
    ai
  • Resonance
    Resonance
    Y Combinator LogoW2024
    Acquired • 2 employees • San Francisco
    Resonance hyper-personalizes MarTech campaign content and automatically refreshes and stores high performing content for re-use.
    reinforcement-learning
    artificial-intelligence
    saas
    subscriptions
    marketing